Multimedia Search: From Relevance to Usefulness
نویسنده
چکیده
T here has been an amazing amount of work done and progress achieved in the field of multimedia search over the past two decades. As nicely elaborated on by Lei Zhang and Yong Rui in their recent review of the advances in this field, the developments so far have gone through three main stages: the text-based stage, the content-based stage, and the Web-based stage. In the text-based stage, multimedia search was essentially realized as a trivial extension of the classic, text-oriented information retrieval approach, where text was found in the documents accompanying multimedia items (images, video, and music). In the 1990s, researchers started to explore the possibilities of analyzing the actual (audio-visual) content of multimedia items to automatically infer semantic similarities and textual labels reflecting what is depicted in an image or a video frame and what is audible in a soundtrack. The main motivation for this stage was that manually adding texts to multimedia items was tedious and time consuming. Such texts were also typically one-sided and therefore not always informative enough to help locate the targeted multimedia item. More recently, inspired by the rapid development of social media platforms, text has come back as an important information source for multimedia indexing and search; however, now it exists in the form of user-generated tags, comments (YouTube), and messages (tweets) accompanying multimedia items being uploaded on such platforms. Although still manually added, this “new text” has become scalable through social interaction. Furthermore, text inserted by various people covers different perspectives, increasing the richness of textual metadata. All this has helped text reenter the field so it can be increasingly exploited by the new wave of approaches marking the Web-based stage, where it is integrated with content-based analysis to make indexing and search more robust and reliable. Different stages have been marked by the dominance of some key theoretical and algorithmic paradigms. For instance, the content-based stage started with an exploration of a broad range of audio-visual signal analysis methods, but it later increasingly turned more toward well-established algorithmic frameworks, such as support vector machines (SVMs). Similarly, graph theory has provided the main analytic framework for the Web-based stage. The most recent algorithmic hype in the field is deep learning. If we analyze these three stages and the related (dominating) technologies, we can conclude that their evolution has essentially been resource-driven. In other words, as we have moved from the text-based stage, via the content-based stage, to the Web-based stage, we have included more and more information resources in the development of methods for indexing and searching multimedia. Such technologies have then been investigated and enhanced to get the most out of these resources and improve the results of multimedia search in view of objective criteria, such as average precision (AP). The key to understanding how a multimedia search system works lies in understanding the notion of relevance. Relevance has so far been explored and optimized mainly with respect to queries. The uncertainty in the channel connecting the query and the collection is typically large, and the need to overcome this uncertainty undoubtedly justifies the tremendous effort invested in the development of various relevance models over the past years, involving more and more information resources and applying new generations of algorithms. An increasing number of results reported in recent literature indicate, however, that pursuing this resource-exploitation target of “technically” optimizing the relevance to a query may have already reached its limits. Specifically, it is unclear how much scientific breakthrough can still be achieved and how much impact all the Alan Hanjalic Delft University of Technology EIC’s Message
منابع مشابه
Surface Features in Video Retrieval
This paper assesses the usefulness of surface features in a multimedia retrieval setting. Surface features describe the metadata or structure of a document rather than the content. We note that the distribution of these features varies across topics. The paper shows how these distributions can be obtained through relevance feedback and how this allows for adaptation of (content-based) search re...
متن کاملAdvertising Keyword Suggestion Using Relevance-Based Language Models from Wikipedia Rich Articles
When emerging technologies such as Search Engine Marketing (SEM) face tasks that require human level intelligence, it is inevitable to use the knowledge repositories to endow the machine with the breadth of knowledge available to humans. Keyword suggestion for search engine advertising is an important problem for sponsored search and SEM that requires a goldmine repository of knowledge. A recen...
متن کاملRelevance Computation for Keyword Searches over Structured Multimedia Metadata
Multimedia search engines rely heavily on content analysis algorithms and knowledge management techniques to adequately populate metadata databases. However once this metadata is available, the search engine must work on its own to extract the greatest possible amount of information from the user’s query and map it to the available metadata structure. This paper presents a relevance computation...
متن کاملThe CUBRIK Project
The Cubrik Project is an Integrated Project of the 7th Framework Programme that aims at contributing to the multimedia search domain by opening the architecture of multimedia search engines to the integration of open source and third party content annotation and query processing components, and by exploiting the contribution of humans and communities in all the phases of multimedia search, from...
متن کاملGlocal Multimedia Retrieval
Personal experience is intrinsically local while common knowledge is global. As a consequence, standard multimedia search engines suffer from a gap between local content and global concept, due to the diversity of context. Here we design a new multimedia retrieval system which integrates local diversity into an evolving global knowledge. A suitable personalization of global ontologies allows to...
متن کاملQuicklook2: An Integrated Multimedia System
The need to retrieve visual information from large image collections is shared by many application domains. This paper describes the main features of the multimedia information retrieval engine of Quicklook2. Quicklook2 allows the user to query image and multimedia databases with the aid of sample images, or an impromptu sketch and/or textual descriptions, and progressively refine the system’s ...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- IEEE MultiMedia
دوره 22 شماره
صفحات -
تاریخ انتشار 2015